Papers with Music AVQA datasets
Learning Sparsity for Effective and Efficient Music Performance Question Answering (2025.acl-short)
Copied to clipboard
| Challenge: | Existing Music AVQA methods rely on dense and unoptimized representations, leading to inefficiencies in the isolation of key information, reduction of redundancy, and prioritization of critical samples. |
| Approach: | They propose a sparse learning framework specifically designed for Music AVQA to address these challenges. |
| Outcome: | The proposed framework reduces training time by 28.32% while maintaining accuracy while maintaining state-of-the-art performance on the Music AVQA datasets. |
Music Audio-Visual Question Answering Requires Specialized Multimodal Designs (2026.findings-acl)
Copied to clipboard
Wenhao You, Xingjian Diao, Wenjun Huang, Chunhui Zhang, Keyi Kong, Weiyi Wu, Chiyu Ma, Zhongyu Ouyang, Tingxuan Wu, Ming Cheng, Soroush Vosoughi, Jiang Gui
| Challenge: | Music audio-visual question answering presents unique challenges with dense audio-visual content, intricate temporal dynamics, and the need for domain-specific knowledge. |
| Approach: | They analyze Music AVQA datasets and analyze their results to identify key design patterns . they propose concrete future directions for incorporating musical priors . |
| Outcome: | The proposed architectures are critical for success in Music AVQA, the authors argue . they suggest concrete future directions for incorporating musical priors . |
Learning Musical Representations for Music Performance Question Answering (2024.findings-emnlp)
Copied to clipboard
| Challenge: | Existing methods for audio-visual learning fail to consider the distinctive characteristics of instruments and music. |
| Approach: | They propose to integrate multimodal interactions within the context of music data and annotate and release rhythmic and music sources in the current music datasets to enable the model to learn music characteristics. |
| Outcome: | The proposed model can learn music characteristics from the current music datasets and align its predictions with the temporal dimension. |